Skip to content

cua_s1: complete native vision and screenshot inference with GPU parity - #64

Open
Levius-Fubuki wants to merge 12 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-native-vision
Open

Levius-Fubuki wants to merge 12 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-native-vision

Conversation

@Levius-Fubuki

@Levius-Fubuki Levius-Fubuki commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Complete native screenshot inference: PNG/JPEG → RGB preprocessing → 24-block CUDA vision encoder and merger → image feature insertion and T/H/W positions → language execution → candidate decisions. omni-cua-s1-vision serves the screenshot contract and reuses image features across questions in each request.

Includes #59, #63 and the refreshed #56. Vision retains BF16 base weights and all 50 FP32 LoRA pairs separately; language uses a merged BF16 export. Startup verifies pinned checkpoint/export identity and hashes, then executes real warmup before readiness.

Refresh against current main's shared Qwen module, optimized kernels and 64-entry graph cache. CPU decoding/tokenization and response reconstruction live in the processor, outside executor admission. Shared SerialScheduler admits one complete request; both CUDA streams drain before release, including failures. Multimodal execution remains eager. CUDA ABI is 5: rebuild both library and worker.

The diff retains runtime code, required dependencies, upstream licenses/notices and the one-time exporter. External tests, reports and controls are kept outside the PR diff.

Build and launch

Obtain the pinned weights/lock and Python environment using recipe/cua_s1/text.md; additionally install torchvision==0.29.0 numpy==2.5.3 Pillow==11.3.0.

PYTHONPATH=src .venv/bin/python recipe/cua_s1/export_multimodal_language.py \
  --base weights/Qwen3.5-4B --adapter weights/cua-s1-4b-0.2/multimodal \
  --out weights/cua-s1-multimodal-language
src/backends/cuda/qwen3_5/build.sh target/release 89
cargo build --release --locked -p omni-cua-s1-native --bins
CUA_S1_BASE=weights/Qwen3.5-4B \
CUA_S1_VISION_ADAPTER=weights/cua-s1-4b-0.2/multimodal \
CUA_S1_MODEL=weights/cua-s1-multimodal-language \
CUA_S1_CUDA_LIB=$PWD/target/release/libqwen3_5_cuda.so \
  target/release/omni-cua-s1-vision

The worker defaults to 127.0.0.1:8000; /health reports readiness after initialization/warmup. Export manifest hashes supply local provenance, not an external signature; keep checkpoint files immutable during execution.

Test Plan

Baseline: f594d7dfc4c2bef812e23f7ed73573be9625b287.
Head: 1152531a704044c879adabf377986d14010edbdb.
Build source: 9538f00c5b3c8eb57b3bbb1c874f32c97b459070, whose tree is byte-identical to final head after merging #56 (f25e47a5bc082c3b94462bc67486f792bbfc3794). Remote production source hashes were verified against the final head.

One RTX 4090 (24,564 MiB), driver 595.71.05, CUDA 13.0.88 / sm89, Rust 1.99.0, Python 3.12.3, PyTorch 2.14.0+cu130, Transformers 5.17.0, PEFT 0.21.0. Base Qwen/Qwen3.5-4B@851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a; multimodal adapter cua-ai/cua-s1-4b-0.2@16818868b0cc7813808aae4e87b417657046ab79. All downloaded files passed pinned size/SHA-256 checks. Vision reference uses separate unmerged BF16/FP32 runs with TF32 disabled; native language uses the newly exported merged BF16 checkpoint.

Frozen standard corpus: 7 requests / 8 questions, 1–26 candidates, PNG/JPEG, portraits/wide images and question reuse. Boundary corpus: 4 requests / 11 questions, including 1×1, 200×1, 1,024×1,024, noise and eight-question requests. Fresh BF16 and FP32 controls were regenerated on this host. Each fixed corpus must satisfy max native probability error <= 2 * max BF16 reference error + 0.01, and match choices where FP32 margin ≥ 0.05. No sample or tolerance changes.

Test Result

Local formatting, strict all-target Clippy, workspace tests and locked release build passed; current-head GitHub Rust, benchmark and docs CI passed (deployment skipped).

  • Six shared CUDA regressions, vision CUDA primitive comparisons, full-model graph reuse/eviction/growth and multimodal boundary tests passed.
  • Standard corpus: 8/8 choices, max error 0.00331324, allowance 0.02052773.
  • Boundary corpus: 11/11 choices, max error 0.07950398, allowance 0.20292068.
  • Token IDs, grids and T/H/W positions matched; repeated language execution and vision replay were identical. PNG decoding matched Pillow; the tested JPEG differs by at most three intensity levels.
  • Three external processor tests passed whole-request validation, ordering/usage and malformed-input rejection before execution.
  • Real worker and frontend proxy each passed 11 requests / 19 questions, matching direct probabilities within 1e-7; nine invalid-input cases passed per corpus.
  • 22 requests at concurrency 8 returned bodies identical to serial execution.
  • Main/cua_s1: accept image embeddings and 3D positions in native language model #56/cua_s1: complete native vision and screenshot inference with GPU parity #64 matched text probes in eager/graph mode. The full 27B Open-Jev checkpoint was not executed; its CPU contracts/workspace checks passed.

Demo / evidence

These are fresh actual GPU/HTTP results for the verified final runtime tree, not the previous cleanup run. Raw output JSON, FP32/BF16 controls, frozen protocol, hashes, setup/build logs and reproduction scripts are retained in artifacts/cua-native-refresh-20261005/ and the matching execution-host directory. Current-head CI. External replay/control tools derive from the pre-cleanup source; updated graph/processor/concurrency harnesses remain local under the core-code constraint.

This finite corpus establishes neither bitwise vision equivalence nor general model accuracy. Latency/throughput and cold-vs-warm performance were not benchmarked; no speed or memory improvement is claimed. Download retries and two PR56 harness/build-target failures are preserved in the evidence. Corrections changed only transfer/test setup, without runtime patches, discarded numerical failures or relaxed tolerances. No video is needed for this API parity run; raw assets are not publicly attached.

Self-review

Full diff and architecture contracts reviewed; independent source review found no actionable findings. Contributor checklist retained below.

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

@Levius-Fubuki
Levius-Fubuki marked this pull request as ready for review October 5, 2026 03:16
Copilot AI balanced review requested due to automatic review settings October 5, 2026 03:16

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants